The Journal of the Acoustical Society of America
● Acoustical Society of America (ASA)
All preprints, ranked by how well they match The Journal of the Acoustical Society of America's content profile, based on 35 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Agarwalla, S.; Farhadi, A.; Carney, L. H.
Show abstract
The role of medial olivocochlear (MOC) efferent gain control in auditory enhancement (AE) was investigated using a subcortical auditory model. AE refers to the influence of a precursor on detectability of targets. The absence (or presence) of a precursor component at the target frequency enhances (or suppresses) detection under simultaneous masking conditions. Furthermore, the enhanced target under simultaneous masking acts as a stronger forward masker for a delayed probe tone, known as AE under forward masking. Psychoacoustic studies of AE report findings that challenge conventional expectations, and the underlying mechanisms remain unclear. For instance, listeners with hearing impairment have AE under simultaneous masking but not forward masking (Kreft et al., 2018; Kreft and Oxenham, 2019), whereas listeners with normal hearing have level-dependent AE under forward masking (Kreft and Oxenham, 2019). Our model with MOC efferent gain control successfully replicated these findings. In contrast, a model without efferent gain control failed to capture these effects, supporting the hypothesis that MOC-mediated cochlear gain modulation may play a role in AE and its alteration by hearing loss.
Umadi, R.
Show abstract
Constant-frequency (CF) bats exhibit rapid oscillations of their external ears. Yet, the functional role of these movements has remained unresolved since their initial documentation over half a century ago. Although recent studies have demonstrated that pinna motion generates Doppler shifts, they do not explain why ear oscillations intensify at close range or how these dynamics contribute to echo perception. In this study, I investigate the hypothesis that oscillatory ear movements enhance echo information during CF echolocation. Using a simplified receiver-motion model, I examine how time-varying pinna pose reshapes the temporal and spectral structure of returning echoes. I show that ear oscillations inject dynamic transformations into the received signal, producing multiple informative views of the same echo and increasing both temporal contrast and spectral diversity around the CF carrier. These transformations are strongest under behavioural conditions in which target-state uncertainty is expected to be high, offering a potential functional explanation for the long-standing observation that ear-oscillation rate increases as bats approach a target. The results suggest that oscillatory ear movements act as an adaptive, receiver-side mechanism that enhances echo information during CF echolocation, complementing the well-known emitter-side adaptations of high-duty-cycle biosonar.
Adjekum, R. N.; Small, S. A.; Chan, S.; Stapells, D. R.
Show abstract
ObjectiveThe current study examined the frequency specificity of NB chirps by comparing the spectral characteristics of 500-, 1000-, 2000- and 4000-Hz NB CE-Chirp(R) LS stimuli with those of 2-1-2 tones. DesignSpectral characteristics including the centre frequency, bandwidth, and stimulus energy changes after stopband filtering were compared. The bandwidth was computed as the difference between the upper and lower frequencies at -20 dB (& -3 dB) cutoff points of the main lobe; the centre frequency was determined as the geometric mean of the upper and lower frequencies at the -20 dB (& -3 dB) cutoff points. ResultsAt 100 dB peSPL, the bandwidths of the 500-, 1000-, and 2000-Hz NB CE-Chirp(R) LS acoustic spectra were 1.7-2.5 times wider than the acoustic spectra for the 2-1-2 tones; the 4000-Hz NB CE-Chirp(R) LS bandwidths were 1.4-1.6 times wider than those of the 2-1-2 tones. The energy of NB CE-Chirp(R) LS stimuli was concentrated within {+/-}0.75 octave of the centre frequency, compared to {+/-}0.5 octave for 2-1-2 tones. ConclusionNB CE-Chirp(R) LS stimuli demonstrated poorer frequency specificity compared with 2-1-2 tones. Further studies are needed to investigate the place specificity of the ABRs to NB CE-Chirp(R) LS before implementing them clinically.
Petersen, E.; Shen, Y.
Show abstract
The Auditory Brainstem Response (ABR) can be used to evaluate hearing sensitivity of animals who are unable to respond to behavioral tasks. However, typical data collection methods are time consuming; decreasing the measurement time may save resources or allow researchers to spend more time on other tasks. Here, an adaptive algorithm is proposed for efficient estimation of ABR thresholds. The algorithm relies on the online update of the predicted hearing threshold from a Gaussian process model as ABR data are collected using iteratively optimized stimuli. To validate the algorithm, ABR threshold estimation is simulated by adaptively sub-sampling pre-collected ABR datasets for which the stimuli were systematically varied in frequency and level. The simulated experiment is performed on 5 datasets of Mouse (2 different datasets), Budgerigar, Gerbil, and Guinea Pig ABRs collected by different laboratories, with a total of 27 ears. The original datasets contain between 68 and 106 stimuli conditions, while the adaptive algorithm is run up to a total of 20 stimuli conditions. The adaptive algorithm ABR threshold estimate is compared against human rater estimates who view the full ABR dataset. The adaptive algorithm threshold matches the human estimates within 10 dB, averaged over frequency, for 19 out of 27 ears. The adaptive procedure is able to provide threshold estimates that are comparable to the human rater estimated thresholds while reducing the measurement time by a factor of 3 to 5. The standard deviation of threshold estimates from successive runs is smaller than the inter-human rater differences, indicating adequate test/retest reliability.
Sinha, R.; Azadpour, M.
Show abstract
Vocoder simulations have played a crucial role in the development of sound coding and speech processing techniques for auditory implant devices. Vocoders have been extensively used to model the effects of implant signal processing as well as individual anatomy and physiology on speech perception of implant users. Traditionally, such simulations have been conducted on human subjects, which can be time-consuming and costly. In addition, perception of vocoded speech varies significantly across individual subjects, and can be significantly affected by small amounts of familiarization or exposure to vocoded sounds. In this study, we propose a novel method that differs from traditional vocoder studies. Rather than using actual human participants, we use a speech recognition model to examine the influence of vocoder-simulated cochlear implant processing on speech perception. We used the OpenAI Whisper, a recently developed advanced open-source deep learning speech recognition model. The Whisper models performance was evaluated on vocoded words and sentences in both quiet and noisy conditions with respect to several vocoder parameters such as number of spectral bands, input frequency range, envelope cut-off frequency, envelope dynamic range, and number of discriminable envelope steps. Our results indicate that the Whisper model exhibited human-like robustness to vocoder simulations, with performance closely mirroring that of human subjects in response to modifications in vocoder parameters. Furthermore, this proposed method has the advantage of being far less expensive and quicker than traditional human studies, while also being free from inter-individual variability in learning abilities, cognitive factors, and attentional states. Our study demonstrates the potential of employing advanced deep learning models of speech recognition in auditory prosthesis research.
Pena, J. A. P.; Calvache, C.; Alzamendi, G. A.; Ibarra, E.; Solaque, L.; Peterson, S. D.; Zanartu, M.
Show abstract
Many voice disorders are linked to imbalanced muscle activity and known to exhibit asymmetric vocal fold vibration. However, the relation between imbalanced muscle activation and asymmetric vocal fold vibration is not well understood. This study introduces an asymmetric triangular body-cover model of the vocal folds, controlled by the activation of intrinsic laryngeal muscles, to investigate the effects of muscle imbalance on vocal fold oscillation. Various scenarios were considered, encompassing imbalance in individual muscles and muscle pairs, as well as accounting for asymmetry in lumped element parameters. The results highlight the antagonistic effect between the thyroarytenoid and cricothyroid muscles on the elastic and mass components of the vocal folds, as well as the impact on the vocal process from the imbalance in the lateral cricoarytenoid and interarytenoid adductor muscles. Measurements of amplitude and phase asymmetry were employed to emulate the oscillatory behavior of two pathological cases: unilateral paralysis and muscle tension dysphonia. The resulting simulations exhibit muscle imbalance consistent with expectations in the composition of these voice disorders, yielding asymmetries exceeding 30% for paralysis and below 5% for dysphonia. This underscores the versatility of muscle imbalance in representing phonatory scenarios and its potential for characterizing asymmetry in vocal fold vibration.
Liu, G. S.; Ali, N.-E.-S.; O Maoileidigh, D.
Show abstract
The neural response of the brainstem to brief sounds, known as the auditory brainstem response (ABR), is widely employed in the laboratory and the clinic to diagnose hearing loss. In contrast to behavioral methods that assess hearing using responses to sounds on a trial-by-trial basis, current ABR approaches are limited to analyzing the average ABR over hundreds of trials. Historically, trial-by-trial ABR analysis has not been possible owing to each trials small signal-to-noise ratio. Here we overcome this limitation and show how to classify individual ABR trials as detected or undetected. We use the distribution of single-trial ABRs to assess supra-threshold hearing and to define psychophysics-like thresholds, which we call auditory brainstem detection (ABD) thresholds. ABD thresholds decrease as more of the ABR epoch is taken into account, whereas traditional ABR thresholds do not change. Above the ABD thresholds and below 90 dB SPL, signal detection is significantly improved by utilizing more of the ABR epoch. Our method also allows us to rank the supra-threshold hearing ability of individual subjects. Despite having normal ABR thresholds, some subjects appear to have supra-threshold hearing deficits. The trial-by-trial method demonstrates that signal detection by the ensemble of auditory neurons in the brainstem is intrinsically stochastic not only at low stimulus levels, but also at levels up to 100 dB SPL. Significance StatementNeural responses to sound can be measured by electrodes placed on a subjects head and are commonly used in the laboratory and the clinic to assess hearing. Although the auditory system must distinguish each sound stimulus from intrinsic noise, current methods for ana-lyzing the response of the brainstem to sound only utilize the average response to hundreds of stimuli. Here we overcome this constraint by showing how to classify an individual sound stimulus as detected or undetected based on each auditory brainstem response. This ap-proach can assess hearing at all stimulus levels, indicates that subjects with normal hearing thresholds can exhibit supra-threshold hearing loss, and potentially extends the types of hearing deficits that can be diagnosed using auditory evoked potentials.
Wong, K. H.; Strimbu, C. E.; Olson, E. S.
Show abstract
Optical coherence tomography (OCT) has allowed in vivo recording of sound-induced vibrations of different regions within the organ of Corti complex (OCC), including the basilar membrane (BM), outer hair cell/Deiters cell (OHC/DC) region, and reticular lamina (RL). In the hook region of the gerbil cochlea, where measurements can be made with a substantially transverse optical axis, the three regions have different and characteristic motion responses: The OHC/DC region has greater motions than the other two regions at frequencies below the best frequency (sub-BF); the RL region typically has the greatest BF peak and smallest sub-BF motion. The phase of the OHC/DC-region motion increasingly lags BM motion phase as frequency increases; the RL-region motion phase leads BM, but with a relatively small value. All three regions are compressively nonlinear in the BF peak, but only the OHC/DC region shows sub-BF compressive nonlinearity. In this paper, we describe the strain that exists within the RL and OHC-body regions. These strains are large where the motion varies over short distances, and a region of large strain can be as short as a single 2.7 {micro}m measurement pixel, or extend over several pixels, with the extensive strains appearing more often at 70 than at 50 dB SPL. Beyond the region of large strain, over a distance that can exceed 20 m, the OHC/DC region displays nearly unvarying motion spatially -- this region appears to vibrate as a body. Statement of SignificanceThe sensory tissue of the cochlea responds actively to a sound stimulus: cell-based forces amplify and enhance the vibration of the sensory tissue. Measurements employing optical coherence tomography have identified major vibration patterns along a sensory-tissue-spanning line that includes the active outer hair cells. In this article, we describe the transitional motion between these major vibration regions and the motion strains that exist as vibration morphs from one region to the next. The findings are presented in frequency response curves to convey the frequency tuning and its stimulus-level dependence, and in one-dimensional heat maps to convey the extent of regional motions and strains. These findings fuel and constrain conceptual and physics-based models of cochlear amplification.
Tubelli, A.; Motallebzadeh, H.; Guinan, J. J.; Puria, S.
Show abstract
A common assumption about the cochlea is that the local characteristic frequency (CF) is determined by a local resonance of basilar-membrane (BM) stiffness with the mass of the organ-of-Corti (OoC) and entrained fluid. We modeled the cochlea while avoiding such a priori assumptions by using a finite-element model of a 20-m-thick cross-sectional slice of the middle turn of a passive gerbil cochlea. The model had anatomically accurate structural details with physiologically appropriate material properties and interactions between the fluid spaces and solid OoC structures. The longitudinally-facing sides of the slice had a phase difference that mimicked the traveling-wave wavelength at the location of the slice by using Floquet boundary conditions. A paired volume-velocity drive was applied in the scalae at the top and bottom of the slice with the amplitudes adjusted to mimic experimental BM motion. The development of this computationally efficient model with detailed anatomical structures is a key innovation of this work. The resulting OoC motion was greatest in the transverse direction, stereocilia-tip deflections were greatest in the radial direction and longitudinal motion was small in OoC tissue but became large in the sulcus at high frequencies. If the source velocity and wavelength were held constant across frequency, the OoC motion was almost flat across frequency, i.e., the slice showed no local resonance. A model with the source velocity held constant and the wavelength varied realistically across frequency, produced a low-pass frequency response. These results indicate that tuning in the gerbil middle turn is not produced by a resonance due to local OoC mechanical properties, but rather is produced by the characteristics of the traveling wave, manifested in the driving pressure and wavelength. STATEMENT OF SIGNIFICANCEThe sensory epithelium of hearing, the organ of Corti, is encased in the bone of the fluid-filled cochlea and is difficult to study experimentally. We provide a new method to study the cochlea: making an anatomically-detailed finite-element model of a small transverse slice of the cochlea using Floquet boundary conditions and incorporating global cochlear properties in the slice drive and the wavelength-frequency relationship. The model shows that the slice properties do not show a mechanical resonance and therefore do not produce the frequency-response tuning of the cochlea. Instead, tuning emerges from global cochlear properties carried by the traveling wave.
Strimbu, C. E.; Chiriboga, L. A.; Frost, B. L.; Fallah, E.; Olson, E. S.
Show abstract
Auditory sensation is based in nanoscale vibration of the sensory tissue of the cochlea, the organ of Corti complex (OCC). Motion within the OCC is now observable due to optical coherence tomography. In the cochlear base, in response to sound stimulation, the region that includes the electro-motile outer hair cells (OHC) was observed to move with larger amplitude than the basilar membrane (BM) and surrounding regions. The intense motion is based in active cell mechanics, and the region was termed the "hotspot" (Cooper et al., 2018, Nature comm). In addition to this quantitative distinction, the hotspot moved qualitatively differently than the BM, in that its motion scaled nonlinearly with stimulus level at all frequencies, evincing sub-BF activity. Sub-BF activity enhances non-BF motion; thus the frequency tuning of the hotspot was reduced relative to the BM. Regions that did not exhibit sub-BF activity are here defined as the OCC "frame". By this definition the frame includes the BM, the medial and lateral OCC, and most significantly, the reticular lamina (RL). The frame concept groups the majority OCC as a structure that is largely shielded from sub-BF activity. This shielding, and how it is achieved, are key to the active frequency tuning of the cochlea. The observation that the RL does not move actively sub-BF indicates that hair cell stereocilia are not exposed to sub-BF activity. A complex difference analysis reveals the motion of the hotspot relative to the frame.
Guest, D.; Cameron, D. A.; Schwarz, D. M.; Leong, U.-C.; Carney, L. H.
Show abstract
Many sounds contain spectral modulations at multiple scales, but much is still unknown about how such spectral features are represented in the auditory system. One behavioral task that provides insight into this question is profile analysis. In a typical profile-analysis task, listeners are asked to discriminate between a complex tone with equal-amplitude components and a complex tone with a single incremented component. Because listeners can perform profile analysis even when the overall sound level of the stimuli is randomized from interval to interval, this task is thought to be a useful index of relative processing of spectral shape, rather than just sensitivity to absolute level changes. Here, we measured profile analysis across the frequency range in a group of listeners that varied widely in their hearing status. We then modeled the resulting behavioral data by decoding responses to the stimuli from computational models of the auditory nerve and inferior colliculus. We found that both hearing loss at the target frequency and increases in the target frequency were associated with poorer profile-analysis thresholds, and that these results could both be explained as the result of corresponding changes in sensitivity of temporal modulation-sensitive cells at the level of the inferior colliculus. These results suggest that key features of profile-analysis may reflect the limits of central neural tuning to temporal modulations.
Verschooten, E.; Strickland, E. A.; Verhaert, N.; Joris, P. X.
Show abstract
Efferent projections from the brainstem to the inner ear are well-described anatomically and physiologically but their precise function remains debated. The medial olivocochlear (MOC) system and its reflex, the MOCR, have been particularly well studied. In animals, anatomical and physiological data are fine-grained and extensive and suggest an important role for the MOCR in anti-masking e.g. to improve the detection of tones in background noise. Extensive behavioral studies in human support this role, but direct linking of behavioral paradigms to the MOCR is challenging because of the difficulty in obtaining appropriate human neural measures. We developed a new approach in which mass potentials were recorded near the cochlea of normal hearing and awake human volunteers to increase the signal-to-noise (SNR) ratio, and examined whether broadband noise to the contralateral ear elicited MOCR anti-masking effects as reported in animals. Probing the mass potential to the onset of brief tones at 4 and 6 kHz, convincing anti-masking or suppressive effects consistent with the MOCR were not detected. We then changed the recording technique to examine the neural phase-locked contribution to the mass potential in response to long, low-frequency tones, and found that contralateral sound suppressed neural responses in a systematic and progressive manner. We followed up with psychophysical experiments in which we found that contralateral noise elevated detection threshold for tones up to 4 kHz. Our study provides a new way to study efferent effects in the human peripheral auditory system and shows that contralateral efferent effects are biased towards low frequencies.
Vesterholm, K. K.; Häfele, F. T.; Figeac, F.; Jakobsen, L.
Show abstract
O_LIAnimals with specialized hearing such as bats utilize the directionality of their hearing for complicated tasks such as navigation and foraging. The directionality of hearing can be described through the head related transfer function (HRTF). Current state of the art for obtaining the HRTF involves either direct measurement with a microphone at the eardrum, or a CT (micro computed tomography) scan to create a 3D model of the head for acoustic modelling. Both methods usually involve dead animals. C_LIO_LIWe developed a 3D photogrammetry approach to create scaled 3D models of bats with sufficient detail to simulate the HRTF using the boundary element method (BEM). We designed a setup of 28 cameras to obtain 3D models and HRTF from live awake bats. We directly compare the mesh models generated by our photogrammetry method and from CT scans as well as the simulated HRTFs from both with measurements using an in-ear microphone. C_LIO_LIGeometries of the mesh models match well between photogrammetry and CT, but with increasing errors where line of sight is compromised for photogrammetry. The resulting HRTFs are in great agreement when comparing CT and in-ear measurements to photogrammetry (correlation coefficients above 0.6). The 3D model and simulated HRTF of the live and awake bat likewise aligns well to the results from the deceased animals. C_LIO_LIPhotogrammetry is a viable alternative to CT scans for the generation of surface models of small animals. These models allow numerical modelling of HRTFs at biologically relevant frequencies. Moreover, photogrammetry allows for model generation and subsequent HRTF simulation of live, awake animals, abolishing the need for euthanasia and anesthesia. It paves the way for large scale acquisition of 3D models for various purposes including HRTFs. C_LI
Azevedo, A. C. P. F. O. d.; Pellegrino, T. G.; Pena, J. L.; Marin, B.; Pavao, R.
Show abstract
Experiments on human auditory perception have shown that interaural time difference (ITD) is sufficient to generate spatial percepts, even though stimuli containing only the ITD cue are perceived as being emitted from inside the head instead of from external locations at specific azimuths. These experiments are thus interpreted as "lateralization" instead of "localization" tasks. In fact, lateralized spatial perception has been quantified using tasks in which participants have to report their estimates by selecting a putative location inside the head, or matching the perceived position to sounds with a given interaural level difference. Therefore, these estimates are made with respect to internal frames of reference, but it is unclear whether these percepts have any significance for the more ecological problem of locating an external sound source. In order to investigate the link between internalized spatial percepts and sound localization, we designed a new task in which subjects are instructed to report externalized azimuthal location for sounds containing only ITD cues. Despite the mismatch between an internalized percept having to be reported as emanating from an external location, subjects were able to estimate azimuths consistently. Furthermore, normalized estimates were indistinguishable from those obtained using traditional lateralization tasks. Our results revealed a direct relationship between perceived azimuths and ITD, which deviates from that obtained from acoustical analysis of binaural recordings, revealing estimation biases. Intriguingly, these results indicate that externalized percepts are not required for the generation of azimuthal percepts.
Gupta, S.; Kalra, L.; Rose, G. J.; Bee, M. A.
Show abstract
Animals often communicate acoustically in noisy social environments, yet how receivers extract relevant information from overlapping sounds is poorly understood. Studies of animal communication in noise are typically based on a traditional filterbank model of hearing focused on energetic masking, where spectrotemporal overlap in peripheral auditory filters limits signal audibility. By contrast, studies of human speech show that background sounds can also interfere with selecting and attending to otherwise audible signals in a phenomenon known as informational masking. Whether informational masking constrains acoustic communication in nonhuman animals remains poorly understood. Through controlled laboratory experiments with treefrogs (Hyla chrysoscelis), we demonstrate, for the first time in a nonhuman animal, that informational masking can impair crucial mate-choice decisions when vocal signals and other concurrent sounds share similar temporal features. These impacts spanned a wide range of signal-to-noise ratios, occurred in the absence of spectrotemporal overlap in the auditory periphery, and were strongest when interfering sounds fell within a frequency band salient for vocal processing. Our findings challenge conventional views on how noise impacts animal communication by establishing informational masking as a general communication problem shared by humans and other animals. These results highlight the need to distinguish between energetic and informational masking to understand the evolution of signaling in complex acoustic environments.
Brughera, A.; Ballestero, J. A.; McAlpine, D.
Show abstract
A potential auditory spatial cue, the envelope interaural time difference (ITDENV) is encoded in the lateral superior olive (LSO) of the brainstem. Here, we explore computationally modeled LSO neurons, in reflecting behavioral sensitivity to ITDENV. Transposed tones (half-wave rectified low-frequency tones, frequency-limited, then multiplying a high-frequency carrier) stimulate a bilateral auditory-periphery model driving each model LSO neuron, where electrical membrane impedance low-pass filters the inputs driven by amplitude-modulated sound, limiting the upper modulation rate for ITDENV sensitivity. Just-noticeable differences in ITDENV for model LSO neuronal populations, each distinct to reflect the LSO range in membrane frequency response, collectively reproduce the largest variation in ITDENV sensitivity across human listeners. At each stimulus carrier frequency (4-10 kHz) and modulation rate (32-800 Hz), the top-performing model population generally reflects top-range human performance. Model neurons of each speed are the top performers for a particular range of modulation rate. Off-frequency listening extends model ITDENV sensitivity above 500-Hz modulation, as sensitivity decreases with increasing modulation rate. With increasing carrier frequency, the combination of decreased top membrane speed and decreased number of model neurons capture decreasing human sensitivity to ITDENV.
Oh, Y.; Hartling, C. L.; Srinivasan, N. K.; Eddolls, M.; Diedesch, A. C.; Gallun, F. J.; Reiss, L. A. J.
Show abstract
In the normal auditory system, central auditory neurons are sharply tuned to the same frequency ranges for each ear. This precise tuning is mirrored behaviorally as the binaural fusion of tones evoking similar pitches across ears. In contrast, hearing-impaired listeners exhibit abnormally broad tuning of binaural pitch fusion, fusing sounds with pitches differing by up to 3-4 octaves across ears into a single object. Here we present evidence that such broad fusion may similarly impair the segregation and recognition of speech based on voice pitch differences in a cocktail party environment. Speech recognition performance in a multi-talker environment was measured in four groups of adult subjects: normal-hearing (NH) listeners and hearing-impaired listeners with bilateral hearing aids (HAs), bimodal cochlear implant (CI) worn with a contralateral HA, or bilateral CIs. Performance was measured as the threshold target-to-masker ratio needed to understand a target talker in the presence of masker talkers either co-located or symmetrically spatially separated from the target. Binaural pitch fusion was also measured. Voice pitch differences between target and masker talkers improved speech recognition performance for the NH, bilateral HA, and bimodal CI groups, but not the bilateral CI group. Spatial separation only improved performance for the NH group, indicating an inability of the hearing-impaired groups to benefit from spatial release from masking. A moderate to strong negative correlation was observed between the benefit from voice pitch differences and the breadth of binaural pitch fusion in all groups except the bilateral CI group in the co-located spatial condition. Hence, tuning of binaural pitch fusion predicts the ability to segregate voices based on pitch when acoustic cues are available. The findings suggest that obligatory binaural fusion, with a concomitant loss of information from individual streams, may occur at a level of processing before auditory object formation and segregation.
Song, R.; Sjons, J.; Ekstrom, A. G.
Show abstract
We present a pipeline for deep neural network assisted modeling and analysis of tube vocal tract models. Such models are composed of a series of cylindrical tube segments, each characterized by length and cross-sectional area. A large synthetic dataset of such tube configurations is generated, and a circuit theory-based algorithm predicts corresponding formant frequencies. To explore the mapping between tube sequence shapes and predicted resonance (formant) values, the pipeline integrates both linear regression and nonlinear machine learning models --including multi-layer perceptrons. Model interpretability is assessed using Shapley Additive Explanations (SHAP), which quantifies the contribution of each segment to predicted formant frequencies. The proposed framework enables detailed exploration of the articulatory-acoustic relationships inherent to an acoustic tube and vocal tract simulacrum. We present and describe the pipeline in the context of modeling effects of perturbations on the first three predicted resonances for a 16- cm tube, divided into 1 cm segments. Our pipeline can be applied to any method that models predictions of behavior of an acoustic tube, where the tube is conceived as a series of segmented units.
Lukashkin, A. N.; Russell, I. J.; Rybdylova, O.
Show abstract
Sensory hair cells, including the sensorimotor outer hair cells, which enable the sensitive, sharply tuned responses of the mammalian cochlea, are excited by radial shear between the organ of Corti and the overlying tectorial membrane. It is not currently possible to measure directly in vivo mechanical responses in the narrow cleft between the tectorial membrane and organ of Corti over a wide range of stimulus frequencies and intensities. The mechanical responses can, however, be derived by measuring hair cell receptor potentials. We demonstrate that the seemingly complex frequency and intensity dependent behaviour of outer hair cell receptor potentials could be qualitatively explained by a two-degrees of freedom system with a local cochlear partition and tectorial membrane resonances strongly coupled by the outer hair cell stereocilia. A local minimum in the receptor potential below the characteristic frequency is always observed at the tectorial membrane resonance frequency which, however, might shift with stimulus intensity.
Simmons, A. M.; Tuninetti, A.; Yeoh, B. M.; Simmons, J. A.
Show abstract
We introduce two EEG techniques, one based on conventional monopolar electrodes and one based on a novel tripolar electrode, to record for the first time auditory brainstem responses (ABRs) from the scalp of unanesthetized, unrestrained big brown bats. Stimuli were frequency-modulated (FM) sweeps varying in sweep direction, sweep duration, and harmonic structure. As expected from previous invasive ABR recordings, upward-sweeping FM signals evoked larger amplitude responses (peak-to-trough amplitude in the latency range of 3-5 ms post-stimulus onset) than downward-sweeping FM signals. Scalp-recorded responses displayed amplitudelatency trading effects as expected from invasive recordings. These two findings validate the reliability of our noninvasive recording techniques. The feasibility of recording noninvasively in unanesthetized, unrestrained bats will energize future research uncovering electrophysological signatures of perceptual and cognitive processing of biosonar signals in these animals, and allows for better comparison with ABR data from echolocating cetaceans, where invasive experiments are heavily restricted. Because experiments can be repeated in the same animal over time without confounds of stress or anesthesia, our technique requires fewer captures of wild bats, thus helping to preserve natural populations and addressing the goal of reducing animal numbers used for research purposes.